human-labelled data
AI's dirty little secret
San Francisco - There's a dirty little secret about artificial intelligence: It's powered by an army of real people. From makeup artists in Venezuela to women in conservative parts of India, people around the world are doing the digital equivalent of needlework - drawing boxes around cars in street photos, tagging images, and transcribing snatches of speech that computers can't quite make out. Such data feeds directly into "machine learning" algorithms that help self-driving cars wind through traffic and let Alexa figure out that you want the lights on. These repetitive tasks pay pennies apiece. But in bulk, this work can offer a decent wage in many parts of the world - even in the US.
Wonder the taskforce behind AI? It's humans
There's a dirty little secret about artificial intelligence: It's powered by hundreds of thousands of real people. From makeup artists in Venezuela to women in conservative parts of India, people around the world are doing the digital equivalent of needlework -- drawing boxes around cars in street photos, tagging images, and transcribing snatches of speech that computers can't quite make out. Such data feeds directly into "machine learning" algorithms that help self-driving cars wind through traffic and let Alexa figure out that you want the lights on. These repetitive tasks pay pennies apiece. But in bulk, this work can offer a decent wage in many parts of the world -- even in the US.
Weakly Supervised PLDA Training
Li, Lantian, Chen, Yixiang, Wang, Dong, Zhao, Chenghui
PLDA is a popular normalization approach for the i-vector model, and it has delivered state-of-the-art performance in speaker verification. However, PLDA training requires a large amount of labelled development data, which is highly expensive in most cases. We present a cheap PLDA training approach, which assumes that speakers in the same session can be easily separated, and speakers in different sessions are simply different. This results in `weak labels' which are not fully accurate but cheap, leading to a weak PLDA training. Our experimental results on real-life large-scale telephony customer service achieves demonstrated that the weak training can offer good performance when human-labelled data are limited. More interestingly, the weak training can be employed as a discriminative adaptation approach, which is more efficient than the prevailing unsupervised method when human-labelled data are insufficient.